Back

IEEE/ACM Transactions on Computational Biology and Bioinformatics

Institute of Electrical and Electronics Engineers (IEEE)

Preprints posted in the last 30 days, ranked by how well they match IEEE/ACM Transactions on Computational Biology and Bioinformatics's content profile, based on 38 papers previously published here. The average preprint has a 0.04% match score for this journal, so anything above that is already an above-average fit.

1
Detecting CYP2C19 deletions from genotyping array signals using neural networks

Yelmen, B.; Hofmeister, R. J.; Lutsar, V. K.; Finianos, M.; Stone, B. C.; Joeloo, M.; Krebs, K.; Kivistik, P. A.; Smit, S.; Estonian Biobank Research Team, ; Metspalu, M.; Hudjashov, G.; Milani, L.

2026-08-25 bioinformatics 10.64898/2026.08.21.746170 medRxiv
Top 0.2%
3.3%
Show abstract

Since copy number variations (CNVs) in pharmacogenes can cause significant alterations in drug metabolism, their reliable detection is of high importance both for large-scale studies and personalized medicine. Whole-genome sequencing, and specifically long-read sequencing, is the gold standard for CNV detection. Despite increasing availability of these technologies, genotyping arrays are still widely used as cost-effective alternatives in biobank and clinical settings, yet calling CNVs based on array intensity signals is challenging due to low base pair resolution. In this work, we developed a neural network model, nnCNV, to predict deletions in the CYP2C19 pharmacogene region from array intensity signals. We compared our method to the most widely used algorithm, PennCNV, and demonstrated better performance reaching 100% accuracy in the test dataset. Furthermore, we predicted probe-by-probe CYP2C19 deletion coordinates for all Estonian Biobank samples using nnCNV and PennCNV, and validated these predictions using an identity-by-descent (IBD) sharing method, which also demonstrated superior nnCNV performance. For the deletion samples with conflicting PennCNV and nnCNV predictions, we performed PCR analysis for validation, which showed 97% precision for nnCNV compared to 23% for PennCNV. Finally, we assessed the gradient-based feature importance maps and showed that nnCNV utilizes signal intensity information not only from deletion probes, but also from probes in flanking regions. Our results demonstrate that long-range information, which cannot be utilized by hidden Markov models, can improve CNV calling.

2
Benchmarking Graph Neural Networks for Multi-Omics Cancer Subtyping using Methylation and Gene Expression Profiles

Schirmacher, J.; Maurer, M. C.; Metsch, J. M.; Ploesch, S.; Chereda, H.; Blumenthal, D. B.; Hauschild, A.-C.

2026-08-25 bioinformatics 10.64898/2026.08.21.745839 medRxiv
Top 0.3%
3.3%
Show abstract

Motivation: Graph Neural Networks (GNNs) have gained increasing interest in the biomedical domain, as the integration of prior knowledge and deep neural networks has the potential to enhance insights into molecular processes and disease mechanisms. However, a comprehensive and systematic assessment of model architectures, data modalities, graph structures, and their performance for graph signal classification in the biomedical domain is yet to be performed. In order to close this gap, we conducted a benchmarking study on multiple GNNs on a Protein-Protein Interaction (PPI) network for Kidney Renal Clear Cell Carcinoma and Breast cancer subtype prediction, performing an in-depth investigation of architectures, incorporating skip connections and various data modalities. Results: While none of the GNNs outperforms the structure-agnostic Multi-Layer Perceptron baseline, all of them can handle bimodal data (gene methylation and expression) and offer the ability to gain explainability based on PPIs. We offer practical guidelines for applying GNNs to graph signal processing tasks specifically for cancer classification. Depending on the underlying dataset and PPI structure employed, models on different data modalities outperform others. Overall, we suggest using ChebNet, which tends to outperform the Graph Convolutional Network and the Graph Attention Network in cancer subtype prediction. We recommend using GNN architectures that employ a simple flattening readout layer, as they provide better classification performance and faster training time than those with global average pooling. Additionally, we tested residual connections, but they had only an insignificant impact on classification performance.

3
A Semantic + Neuronal Approach to Predict Pathogenic Variants in DNA Sequences

Motta, J. A.; Motta, M. d. M.; Fernandez, C.

2026-08-20 bioinformatics 10.64898/2026.08.16.745093 medRxiv
Top 0.3%
3.2%
Show abstract

In this work, we present a machine learning model for identifying pathogenic DNA variants. The model was learned from the analysis of normal and pathogenic sequences extracted from the ClinVar database (supported by NCBI). This analysis was based on a conceptual semantic model of DNA sequences converted to peptide sequences (amino acid sequences) governed by a well-defined grammar, which allowed us to apply NLP techniques, specifically Part of Speech tagging (POS tagging). Our predictive model was built by combining two techniques: CRF (from the Markov model family), which performs the sequencing, and BiLSTM (a deep learning model) which captures the past and future content of the sequences. The training space was created with the sequences of 105 genes associated with approximately 27,000 pathogenic variants. The model was evaluated using the metrics precision, P-R and ROC curves, AUC, and confusion matrices. Its performance was also compared against five known methods for predicting pathogenic variants. The results show exceptional performance that exceeds expectations and places this new method at the state of the art for predicting pathogenic DNA sequences.

4
Genetic Architecture and Sample Size Impact Relative Performance of Nonlinear Machine Learning and Standard Polygenic Risk Scores

Zhu, J.; Baousi, A.; Morris, A. P.; Guo, H.

2026-09-03 genetic and genomic medicine 10.64898/2026.08.29.26361109 medRxiv
Top 0.4%
2.4%
Show abstract

Standard polygenic risk scores (PRSs) are constructed based on additive genome-wide association study (GWAS) summary statistics. Nonlinear machine learning methods have been increasingly applied to construct PRSs directly from individual-level data, with the aim of improving predictive performance over standard PRSs through their ability to model non-additive genetic effects. However, their superiority across studies has been inconsistent, and the conditions under which they provide meaningful improvements remain unclear. We combined theoretical analysis, simulations and a real-world application to investigate when two widely used nonlinear machine learning methods, random forest and XGBoost, outperform standard PRSs. Theoretical analysis showed that standard PRSs can implicitly capture part of the genetic variance attributable to nonadditive genetic effects through their contributions to marginal SNP effects, thereby losing less information than commonly assumed. Although nonlinear models have a higher theoretical potential, their greater flexibility incurs a bias-variance trade-off that can limit predictive gains at finite sample sizes. Simulations showed that XGBoost outperformed the standard PRS only when the genetic architecture involves a sufficiently large proportion of interaction genetic variance concentrated across relatively few interaction effects and large training samples were available. Random forest consistently underperformed the standard PRS. In an application to ischemic heart disease prediction using UK Biobank data, XGBoost showed no meaningful improvement in predictive performance over the standard PRS, whereas random forest again performed worse. Together, these findings suggest that nonlinear machine learning do not uniformly outperform standard PRSs; rather, their relative performance depends jointly on genetic architecture and training sample size. Our study helps to reconcile the inconsistent results reported across previous studies and provides a framework for identifying settings in which more complex PRS models are likely to be beneficial.

5
Relational Graph Convolutional Networks for Glioblastoma Biomarker Discovery via ceRNA and Copy Number Variation Analysis

Khandelwal, S.; Jarvis, N.; Zhan, J.

2026-08-20 bioinformatics 10.64898/2026.08.16.744525 medRxiv
Top 0.4%
2.1%
Show abstract

Glioblastoma (GBM) is a highly aggressive brain tumor with an extremely poor 5-year survival rate of 6.9%, largely attributable to the lack of reliable biomarkers. While competing endogenous RNA (ceRNA) and copy number variation (CNV) analyses offer unique biomarker identification potential, current approaches neglect the integration of multiple regulatory mechanisms for biomarker detection. To address this limitation, we applied relational graph convolutional networks (RGCNs) to ceRNA and CNV knowledge graphs through a novel late fusion ensemble architecture. The proposed architecture outperformed baseline models and identified five novel biomarkers, including hsa-miR-196a and hsa-miR-224. Kaplan-Meier survival analysis and Cox regression indicated that the identified genes hold significant prognostic and diagnostic power. The early stratification of the Kaplan-Meier curves indicates the potential these genes hold for patient survival prediction. The results illustrate that a late fusion RGCN ensemble effectively captures complex gene interactions, overcoming limitations of existing models and providing a framework for biomarker discovery. The novel biomarkers serve as prospective targets for future GBM therapeutic development and candidates for non-invasive diagnostic assays.

6
Beyond Chemical Similarity: Structure-Agnostic Drug-Drug Interaction Prediction with MeSH Semantics and a Drug-Target-Protein Knowledge Graph

Yılmaz, A.; Szydlik, S.; Taheri, G.

2026-08-18 bioinformatics 10.64898/2026.08.10.743843 medRxiv
Top 0.6%
1.7%
Show abstract

BackgroundAdverse drug-drug interactions (DDIs) cause preventable hospitalizations, but exhaustive experimental screening of all drug pairs is infeasible. Many computational predictors rely on SMILES or other molecular representations, limiting their direct applicability to biologics and other non-small-molecule therapeutics. We present a structure-agnostic framework that combines semantic representations derived from Medical Subject Headings (MeSH) with graph-derived topology from a Drug-Target-Protein knowledge graph constructed from DrugBank and UniProt. We further investigate how variation in MeSH annotation depth affects predictive performance. ResultsDrugs are grouped according to their deepest MeSH annotation level (Low, Mid, or Deep), and performance is evaluated across the resulting interaction categories in transductive and inductive settings. The Intermediate ontology scope (Low+Mid) provides the most stable performance, while adding Deep-level terms offers limited and inconsistent benefit. Lightweight topological descriptors are integrated with MeSH features through instance-wise, dimension-specific latent-space gating, using curated reliable-negative pairs for supervision. Fusion improves mean performance over the MeSH-only baseline across all six categories in the transductive setting. Under induction, the clearest gains occur for Low-Low interactions ({Delta}AUROC = 0.056;{Delta} F1 = 0.137) and Low-Mid interactions ({Delta}AUROC = 0.077;{Delta} F1 = 0.114). ConclusionsMeSH annotation depth is associated with systematic variation in DDI prediction performance that aggregate evaluation can obscure. Graph-derived topology is particularly beneficial when ontology annotations are shallow. The framework provides a common, structure-agnostic representation compatible with both small-molecule and biologic therapeutics and supports first-pass DDI prioritization for subsequent expert assessment.

7
DQHTFI: Dynamic-Query Hypergraph Transformer for Fine-Grained Drug-Target Interaction and Affinity Prediction

Tao, K.; Chai, H.; Chen, Z.; Gao, X.; Yu, B.

2026-08-14 bioinformatics 10.64898/2026.08.08.743505 medRxiv
Top 0.6%
1.7%
Show abstract

Drug-target interaction prediction and binding affinity prediction are two key tasks in drug discovery and drug repurposing. Although deep learning methods have made significant progress, existing models typically rely on global representations of drugs and proteins, making it difficult to adequately model fine-grained interactions between their local units. Fixed multimodal fusion strategies also struggle to dynamically adjust the contributions of different modalities for different drug-target combinations. To address these issues, we propose DQHTFI, a fine-grained interaction prediction framework for drug-target interaction classification and binding affinity regression. DQHTFI employs BRICS fragments and Pfam functional domains as the basic interaction units and jointly learns semantic and structural representations. We design a dynamic-query hypergraph Transformer framework in which hyperedges are constructed among the multimodal features of fragment-domain pairs. Dynamic queries are generated from the cross-conditioned features of fragment-domain pairs to adaptively adjust the contribution of each modality, thereby modeling higher-order interactions between local units. Our proposed model achieves competitive results on multiple benchmark datasets.

8
ProtFinder: An efficient machine learning framework for protein model selection on real data

Nguyen Huy, T.; Dong, Y.; Ly-Trong, N.; Vinh, L. S.; Minh, B. Q.

2026-08-07 evolutionary biology 10.64898/2026.08.04.742760 medRxiv
Top 0.6%
1.7%
Show abstract

Model selection is a fundamental step in phylogenetic analysis that determines the best-fit model of sequence evolution for a given multiple sequence alignment. Popular model selection methods, such as ModelFinder, rely on statistical information criteria, such as the Bayesian Information Criterion (BIC) or the Akaike Information Criterion (AIC). However, these approaches are computationally expensive and the use of information criteria has been the subject of ongoing discussion. Recently, machine learning has emerged as a promising approach for phylogenetic model selection in both nucleotide and protein sequence analyses. ModelDetector is currently the only machine learning-based method for amino acid substitution model selection. However, because ModelDetector was trained on simulated data, it does not perform well on real datasets. Another limitation is that it does not support different rate heterogeneity across sites (RHAS) models. To overcome these limitations, we introduce ProtFinder, an efficient machine learning framework for protein model selection that predicts amino acid substitution models, RHAS models, and amino acid frequency models. To enable ProtFinder to work with real datasets, we employed a transfer learning strategy consisting of three stages: (1) initial training on large-scale simulated data, (2) joint training on both simulated and real data, and (3) final fine-tuning using real data only. Experimental results show that ProtFinder outperformed ModelDetector in amino acid substitution model selection. ProtFinder achieved comparable accuracy to the maximum likelihood method ModelFinder for substitution model selection on medium and large MSAs. It performs slightly better than ModelFinder in RHAS model selection and substantially outperforms it in amino acid frequency model determination. Notably, ProtFinder is up to 1,400 times faster than ModelFinder in terms of inference time, making it particularly suitable for medium and large datasets.

9
Model Validation Protocols for Machine Learning in Small Molecule Drug Discovery

Seal, S.; Zalte, A. S.; Araripe, D. A.; Gomes, R. A.; Korani, D.; Shekhar, M.; Siramshetty, V. B.; Patra, A.; Mou, Z.; Yu, X.; Kuhn, D.; Weskamp, N.; Ash, J.; Cheng, A. C.; Fang, C.; Price, D.; Aldeghi, M.; Rodriguez-Perez, R.; Clevert, D.-A.; Engkvist, O.; Deibler, K.; Rouquie, D.; Reutlinger, M.; Richmond, N. J.; Ainsley, J.; Ledeboer, M.; Green, W. H.; Bender, A.; Wognum, C.

2026-08-24 bioinformatics 10.64898/2026.08.19.745868 medRxiv
Top 0.6%
1.7%
Show abstract

Machine learning (ML) models for molecular property prediction are increasingly deployed in drug discovery, yet their adoption in real-world scenarios requires an understanding of the conditions in which a model succeeds or fails. While standardized benchmarks are powerful instruments to measure and unlock progress in ML research, they should not be blindly treated as the end goal. Especially static and retrospective benchmarks, in which no true unknown test set is employed, limit our ability to robustly validate a model's performance. Building on the collective expertise of a cross-industry consortium, we present a model validation framework consisting of five recommendations that would enable the community to move beyond aggregate metrics toward understanding where and why molecular property prediction models fail. We connect evaluation choices to real-world applications and case studies encountered in pharmaceutical research. The framework proposes splitting strategies that mimic realistic distribution shifts and expose common failure modes. We apply the recommended framework to a recently released dataset of absorption, distribution, metabolism, and excretion (ADME) properties. Across two complementary model algorithms, our case studies reveal four distinct failure modes (extrapolation, interpolation, representation, and evaluation), showing that model errors arise not only from distribution shift but also from limitations in molecular representations. Our results show that commonly used evaluation protocols can significantly overestimate performance and may not detect important model failure modes. All software and data are released via https://github.com/srijitseal/polaris.

10
Clinically Generalisable End-to-End Graph Learning for CT Image-Based Multitask Stroke Diagnosis

Lu, Z.; Uddin, S.; Uribe, S.; White, S.; Martins, R. T.; Chau, S.; Mosaddek, A. S. M.; Islam, M. S.; Nahar, N.; Azad, A. K. M.; Hossain, K. M. N.; Choudhury, H. S.; Hasan, K. M. R.; Mosaddek, N.; Rahman, S.; Hossain, M. M.; Sizar, K. M. M. H.; Angione, C.; Lio, P.; Islam, M. T.; Moni, M. A.

2026-08-31 radiology and imaging 10.64898/2026.08.26.26360026 medRxiv
Top 0.6%
1.7%
Show abstract

Stroke remains a leading cause of mortality and long-term disability worldwide, yet rapid diagnosis is often limited by the shortage of trained radiologists, particularly in resource-constrained settings. Automated analysis of CT imaging offers a potential solution, but existing methods often struggle to achieve clinically generalisable performance while jointly addressing multiple diagnostic tasks. Here we present the Intelligent Integrated Stroke Diagnosis System IISDS, an end-to-end deep learning framework built upon StrokeGNN, a graph-based architecture that integrates 3D contextual feature extraction with U-Net-based 2D lesion segmentation to enable comprehensive stroke analysis from non-contrast CT scans. IISDS performs stroke subtype classification, lesion segmentation and lesion volume estimation within a unified pipeline. To develop and validate the system, we collected and curated BGD-ISD through a collaboration between AI researchers, neurologists, radiologists and clinicians, resulting in a large multi-centre dataset comprising 1,507 CT scans from 597 stroke cases acquired across six hospitals and medical centres in Bangladesh. Across BGD-ISD and multiple publicly available datasets, IISDS achieves state-of-the-art performance on all tasks, improving segmentation accuracy by [≥]0.011 Dice score, reducing lesion volume estimation error by [≥]0.3 average symmetric surface distance (ASSD), and increasing classification performance by [≥]0.018 area under the receiver operating characteristic curve (AUC) compared with existing approaches. These results demonstrate the potential of graph-based deep learning to enable clinically generalisable, automated and scalable stroke diagnosis from CT imaging, supporting rapid clinical decision-making, particularly in healthcare environments with limited access to expert radiological interpretation.

11
Principal Genes: A PCA-based approach to highly variable genes selection for scRNA-Seq analysis

Kakwambi, E. D.; Nguyen, T.; Kapoor, S.; Moussa, M. R.

2026-08-13 bioinformatics 10.64898/2026.08.07.743504 medRxiv
Top 0.7%
1.4%
Show abstract

Single cell RNA-sequencing (scRNA-Seq) data are typically represented as cell-by-gene count matrices, which capture the expression of each gene as detected in the sampled cells; often a heterogeneous population of multiple different cell types or cell states. Almost all scRNA-Seq analysis workflows have a gene selection step prior to applying clustering algorithms which helps remove genes with low variability and hence reduce the high-dimensional gene space. A de-facto method for achieving selection of highly variable genes (HVG) uses dispersion and mean expression scores to evaluate the variability of each individual gene. However, methods based on direct mean-to-variance relationship for gene selection often suffer from susceptibility to variance instability and arbitrary determination of the optimal number of genes to use in downstream analysis tasks, additionally, they often prioritize genes with low abundance but high variance. Here, we propose an innovative method for selecting highly variable genes that is not based on mean to variance ratios: "Principal Genes (PG)" method; it utilizes the rotations (or loadings) from Principal Component Analysis (PCA) to calculate a novel variability score per gene that we name "Gene Principal Score (GPS)". GPS helps evaluate the genes based on their contribution in the PCA rotations and hence ranks the genes according to their variability from highest to lowest variable genes. For efficient implementation we utilize Augmented Implicitly Restarted Lanczos Bidiagonalization methods to efficiently obtain Principal Components (PCs) associated with the largest variance. Genes with the highest GPS score, i.e. Principal Genes, can then be used for downstream analysis tasks, especially the clustering step. To test the performance of our highly variable gene identification method, we use several validation strategies, including clustering of labeled single cell RNA-Seq data (i.e. data with known ground truth cell type labels). Furthermore, we measure the performance of our method against dispersion-based highly variable gene (HVG) selection approaches. We use several validation metrics, including sensitivity and adjusted rand index scores for clustering based on genes selected using our method against genes selected using HVG; and our validation datasets include six real labeled single cell RNA-Seq datasets. Our findings show that our new method, Principal Genes, is comparable and often favorable in performance in selecting highly variable genes and achieves ultra-fast gene selection from PCA results.

12
Moirai: single-cell trajectory inference grounded in gene-level expression dynamics

Fijn, A. H. B.; S. Jeuken, G.

2026-08-11 bioinformatics 10.64898/2026.08.05.742709 medRxiv
Top 0.8%
1.1%
Show abstract

Underlying the development of multicellular organisms is the process of cell differentiation, which is governed by the concerted and sequential change in gene expression. Various methods have been developed that employ scRNA-seq data to infer the position of a cell along a pseudo-temporal axis and identify relevant genes involved in the process. These trajectory inference methods typically rely on global transcriptomic changes and mathematical methods. However, overemphasis on large-scale transcriptomic changes may impair sensitivity to identify branching points and convergent trajectories, which are rather governed by small-scale transcriptional events. Motivated by this, we developed Moirai, a graph-based trajectory inference method that identifies gene expression patterns that change dynamically over a developmental continuum and leverages these to define a common pseudotime axis between all cells. In doing so, Moirai shifts the focus to individual gene dynamics, which enhances its ability to detect putative branching points that are masked by global transcriptomic similarities. We apply Moirai to four developmental datasets, where we demonstrate its ability to recover gene expression patterns of genes with a known involvement in the respective developmental process, motivating their use for defining a cells pseudotime. We furthermore show that Moirai can robustly infer gene expression patterns across different embedding approaches, highlighting the value of moving the focus of the inference process to the small-scale transcriptional dynamics.

13
A Biologically Informed Heterogeneous Graph Neural Network for Multi-Task Prediction of ncRNA-Metastasis-Cancer Interactions

Midjani, F.; Shaghouzi, M.; Banadaki, A. D.; Rahimikashkooli, N.; Keshtkar, F. Z.; Malekpour, M.; Hashemi, S.; Hernandez-Barco, Y. G.; Soleymanjahi, S.

2026-08-21 systems biology 10.64898/2026.08.18.745571 medRxiv
Top 0.8%
1.1%
Show abstract

Metastasis involves context-dependent molecular interactions in which non-coding RNAs, particularly miRNAs and circRNAs, play important regulatory roles. However, existing computational approaches generally do not jointly represent cancer type, metastatic event, and cancer-specific metastatic context. We developed a context-aware multi-task heterogeneous graph neural network (GNN) for predicting ncRNA associations with cancer types and metastatic events. The framework integrates multiple biological repositories into a heterogeneous graph representing ncRNAs, cancers, metastatic event types (METs), and cancer-specific metastatic instances (CSMIs). The model performs six link-prediction tasks using a hierarchical transformer-based encoder and multi-relational TuckER decoder. Across ten independently initialized runs evaluated on the RNA-group-disjoint held-out test set, the model achieved a global AUROC of 0.8801 {+/-} 0.0118 and an F1 score of 0.8260 {+/-} 0.0071. All three ablation variants yielded lower AUROC, with the largest reduction under independent task training. Case studies in pancreatic cancer, colorectal cancer, and hepatocellular carcinoma provided disease-level, event-level, and expression-based support, respectively, for top-ranked candidate associations. The framework enables context-specific prioritization of ncRNA-cancer-metastasis associations for experimental evaluation.

14
SLIM: A small linear model with STRING embeddings for single-cell genetic perturbation prediction

Hu, D.; Pielies Avelli, M.; Jensen, L. J.; Rasmussen, S.

2026-08-07 bioinformatics 10.64898/2026.08.07.743481 medRxiv
Top 0.9%
1.0%
Show abstract

Predicting cellular responses to genetic perturbations is central to understanding gene function and prioritizing therapeutic targets, but experimental screens cannot exhaustively cover genes, cell types, and perturbation combinations. Recent benchmarks have shown that simple baselines can match or outperform substantially more complex models, suggesting that informative biological priors may be as important as model capacity. Here we present SLIM, a lightweight extension of the bilinear model of Ahlmann-Eltze et al. SLIM represents perturbations with 64-dimensional embeddings derived from the STRING protein network and predicts mean transcriptional responses through a closed-form ridge-regression estimator. It then constructs single-cell populations by retrieving training cells and rescaling each gene to match the predicted mean. We evaluated SLIM against four deep learning models and two simple baselines on four single-gene perturbation datasets and one combinatorial perturbation dataset. Across these within-dataset benchmarks, SLIM achieved competitive mean-response accuracy, ranked first in eight of twelve single-gene dataset-metric comparisons, and produced substantially lower maximum mean discrepancy values than the evaluated alternatives. The model has 640 trainable parameters and fitted each benchmark dataset in under 10 seconds on a CPU. These results show that compact biological representations can support accurate and computationally efficient perturbation prediction. Code is available at https://github.com/RasmussenLab/SLIM. Key PointsO_LISLIM combines a closed-form bilinear predictor with STRING-derived perturbation embeddings. C_LIO_LIAcross five within-dataset benchmarks, SLIM achieved competitive mean-response prediction with only 640 trainable parameters. C_LIO_LISLIM builds cell populations by rescaling retrieved training cells to the predicted mean, so they inherit realistic cell-to-cell variation and gene-gene covariation. C_LIO_LIThe results highlight the importance of perturbation representations and population-construction procedures in low-data benchmarks. C_LIO_LISLIM fits each benchmark dataset in under 10 seconds on a standard CPU. C_LI

15
Causally-inspired meta-representation learning framework for predicting patient-specific clinical responses to drug combinations

Zhang, Q.-Q.; Zhang, S.-W.; Shi, M.-H.; Li, J.-N.; Qiang, Y.-R.; Zhang, T.-H.

2026-08-21 bioinformatics 10.64898/2026.08.13.744613 medRxiv
Top 1%
0.8%
Show abstract

Large-scale prediction and assessment of clinical patient responses (i.e., RECIST class) to drug combinations remains challenging due to scarce patient-derived data. The existing prediction methods mainly rely on cancer cell line models. However, substantial biological heterogeneity between cancer cell lines and cancer patients within same tissues, as well as the heterogeneity between one tissue and another, often limit the generalizability of these methods in clinical patients. To overcome these limitations, here we present CaMeRe, a Causally-inspired Meta-representation learning framework designed to predict patient-specific clinical Response to drug combinations. In situations where stable causal factors and domain-specific response-modulating factors are unobservable, explicit discrete domain labels are unavailable, and data is scarce, CaMeRe designed a domain-invariant causal representation learning (DICRL) model guided by the invariant information bottleneck theory and causal intervention invariance principle, and also built a meta-learning framework with bi-level domain generalization to optimize DICRL model for achieving multi-domain generalization within and across tissues. By integrating the causal representation learning and meta learning framework, CaMeRe not only exhibited robust multi-domain generalization performance across multiple clinical drug combination response datasets and PDXs drug combination response datasets and generalization scenarios, but also had better interpretability. We applied CaMeRe to predict drug-combination response scores for 3,423 patients across 542,080 drug combinations. The predicted scores were significantly associated with biomarkers of known drug combinations and enabled the prioritization of candidate drug combinations across 11 cancer types, with stronger support from literature and clinical trial evidences than random baselines. We believe that CaMeRe can be a useful tool for predicting large-scale clinical individual drug combination responses and it has broad clinical applications.

16
An AI System for Autonomous Algorithm Evolution in Drug Development

Zhou, Z.; Nan, Y.; Mou, M.; Qian, Y.; Liu, Y.; Zuo, Z.; Yang, H.; Xu, W.; Li, B.; Jiang, W.; Ren, Y.; Liao, Y.; Wang, Y.; Li, Y.; Yang, Q.; Xi, Z.; Mi, T.; Sun, H.; Liu, P.; Zhu, F.

2026-08-20 pharmacology and toxicology 10.64898/2026.08.16.745117 medRxiv
Top 1%
0.8%
Show abstract

Artificial intelligence (AI) is increasingly permeating the drug development pipeline. Numerous algorithms for accelerating this multi-stage and multi-task process have been constructed, which depends heavily on expert design and labor-intensive task-specific optimization. Given that AI-driven acceleration of drug development is recognized as a cumulative, often synergistic, effect across multiple stages, the autonomous evolution of existing algorithms across the entire pipeline is demanded to achieve a holistic advancement. Here, we present DrugEvolve, a multi-role large language model system for systematic and autonomous algorithm evolution in drug development. DrugEvolve realizes a closed-loop evolution process by incorporating Researcher, Engineer, and Analyst domains, and enables an iterative design, implementation, evaluation, and refinement of algorithm by leveraging scientific knowledge and accumulated evolutionary experience. Across eleven representative tasks spanning target identification, drug discovery, preclinical study, and clinical trial, DrugEvolve autonomously evolved the corresponding task-specific algorithms and achieved substantial performance enhancement on 120 benchmark test sets. Moreover, it showed robust generalizabilities across heterogeneous data modalities (ranging from biological sequence and graph to molecular topology and textual language), and realized gains in both predictive and generative tasks. Collectively, this AI system can serve not only as an algorithmic infrastructure for drug development, but also as a transferable paradigm for broader scientific domains.

17
Autonomous Spatial Transcriptomics Analysis (ASTA): Demonstrating Performance Improvements through Clustering, Biological Annotation, and AI-Driven Discovery

Zhang, M.; Roe, M.; Pollett, C.; Andreopoulos, W. B.

2026-08-18 bioinformatics 10.64898/2026.08.10.743848 medRxiv
Top 1%
0.8%
Show abstract

Spatial transcriptomics keeps measurement of gene expression while preserving spatial context, yet traditional analysis methods face challenges in computational efficiency, biological interpretability, and autonomous discovery. This project presents a framework solving these issues through three parts: (1) an ensemble clustering system achieving 66.7% improvement over baseline average and 23.9% over best single method with silhouette score of 0.540 and statistical significance (p = 0.0032, Cohens d = 1.82); (2) a knowledge-based clustering framework that annotates 88.6% of cells across 8 ovarian cell types using 428 marker genes; and (3) a GPT-4o-mini-powered autonomous agent that generated 3 biological hypotheses with validations.

18
Assessing Computational Models for Pharmacogenomic Variant Interpretation

Pucci, F.; Hermans, P.; Tsishyn, M.; Cusato, J.; Rooman, M.

2026-08-09 bioinformatics 10.64898/2026.08.03.742561 medRxiv
Top 1%
0.8%
Show abstract

Accurately predicting the effects of pharmacogenomic variants is essential for the development of personalized therapeutic strategies, as genetic variability can influence drug response differently across patients. Here, we assessed several computational approaches using a dataset of pharmacogenomic variants with either clinical annotations or functional characterization by deep mutational scanning, compiled from the literature, with an additional focus on CYP2C9, a clinically relevant drug-metabolizing enzyme. Our results show that, despite recent methodological advances, substantial room for improvement remains. In particular, current methods struggle to distinguish gain-of-function variants associated with increased drug clearance and fast-metabolizer phenotypes from neutral variants, whereas loss-of-function variants that reduce drug clearance are predicted more accurately. The integration of structural and evolutionary information appears to be a key strategy for improving performance, with the coevolution-based StructureDCA method achieving the highest accuracy compared with classical genetic variant-effect predictors and recent deep learning approaches, including the pathogenic-variant predictor AlphaMissense and general protein language model-based methods. Finally, our results indicate that computational models can complement in vitro experiments in clinical variant interpretation, as StructureDCA predictions showed better agreement with clinically annotated phenotypes than large-scale deep mutational scanning data in several cases.

19
Benchmarking Imputation Methods for Single-Cell RNA Sequencing Data Using Peripheral Blood Mononuclear Cells from Acute Myocardial Infarction Patients

Ramesh, P.; Fyta, M.

2026-08-27 bioinformatics 10.64898/2026.08.23.746230 medRxiv
Top 1%
0.8%
Show abstract

Acute myocardial infarction (AMI) remains one of the leading causes of mortality worldwide, and the following post-effects, such as post-AMI inflammation and tissue repair, involve peripheral blood mononuclear cells playing a critical role. The influence of imputation methods in biological data is assessed with respect to high-resolution single-cell RNA sequencing (scRNAseq) data relevant to these cells. Still scRNAseq data often encounter a lot of dropout events, leading to sparse and noisy datasets, hampering downstream results. To assess the influence of the missingness in the data, we artificially impose different levels of dropout in available scRNAseq data by leveraging various imputation techniques. Specifically, we introduce artificial missingness at 10%, 20%, and 30% levels under a missing completely at random (MCAR) framework, repeated across 10 independent runs. We benchmarked six imputation strategies - MAGIC, IterativeImputer, KNNImputer, Mean Imputation, SoftImpute, and a Generative adversarial network (GAN) - based approaches using multiple evaluation metrics: marker gene preservation, clustering consistency (Adjusted Rand Index - ARI), gene-wise correlation with ground truth, and structural separation (silhouette scores). The results clearly underline that no single imputation method dominated across all metrics. Overall, Mean and KNN imputers showed limited recovery across all benchmarks. GAN excelled in global transcriptional recovery and SoftImpute in preserving biologically meaningful cell-type signals. Our results highlight the importance of selecting the imputation methods as part of the pre-processing step towards the downstream biological questions related to transcriptome recovery, detection of marker genes, or maintaining cell-type-specific resolution.

20
Inferring disruption of directed graphs using LIKA reveals altered protein phosphorylation networks in schizophrenia

Zhang, L.; Demarco, A. G.; Ghafari, K.; Devlin, B.; MacDonald, M. L.; Roeder, K.

2026-08-07 bioinformatics 10.64898/2026.08.06.743374 medRxiv
Top 1%
0.6%
Show abstract

MotivationKinases regulate a multitude of protein functions, and their dysregulation is pivotal for many human diseases. Direct measurement of kinase activity, however, is often challenging; therefore, inferring activity from the behavior of their substrates is a widely adopted strategy. Nonetheless, traditional methods typically oversimplify the underlying network, ignoring that any particular substrate can be phosphorylated by multiple kinases. ResultsWe present LIKA, a likelihood-based framework for inferring kinase activity from phosphoproteomic data. By modeling the many-to-many structure of kinase-substrate interactions, LIKA achieves high efficiency, even with limited data, while capturing network complexity. Simulation and cell line analyses confirm the robustness and accuracy of LIKA. Importantly, analysis of a phosphoproteomic dataset from schizophrenia and control subjects reveals novel dysregulated kinases. Availability and ImplementationThe implementation code and publicly available data are provided at: https://github.com/lujingz/LIKA.